01 The Big Picture
An LLM API call is stateless, single-shot, and blind: it sees exactly the tokens you send and does exactly one pass. Everything that makes an agent — persistence, tools, planning, memory, permissions — is software wrapped around that call. That software is the harness.
Doc 09 catalogued the parts: skills, rules, workflows, memory banks, subagents. Doc 12 animated the tool-use loop. This doc zooms out to the architecture level and asks the engineering question: given the hardware constraints, why are harnesses shaped the way they are? The answer keeps coming back to three walls established in docs 06–09:
Tokens are currency
Every token in the prompt is re-sent and re-priced every turn (doc 07). A harness that lets context grow linearly with the task pays a compounding bill.
Cache economics rule structure
Cached prefix ≈ 0.1× cost. Harnesses are literally shaped so the expensive parts stay prefix-stable (doc 07) — that's why skills load on demand instead of sitting in the system prompt.
Long context is slow and dilute
Every decode step reads the whole KV cache (docs 06, 08, 22), and attention quality degrades on long transcripts (doc 19, lost-in-the-middle). Small per-call context isn't just cheaper — it's smarter.
Keep the mapping discipline from doc 09: for every mechanism below, we name the bottleneck it relieves. A harness feature that relieves no bottleneck is decoration.
02 What a Harness Is — Precisely
A harness is everything around the model that (a) assembles the context sent to it and (b) executes the loop that turns its outputs into actions. The model proposes; the harness disposes — it runs tools, spawns workers, writes memory, enforces permissions, and decides what the model sees next turn. Production harnesses (Claude Code is the canonical example) decompose into six layers. Read top to bottom as "what the model sees" → "what runs around it":
Layer 1 is inside the model; layers 2–6 are the harness proper. The single-agent context loop — system prompt + tool schemas + memory files + transcript, re-sent every turn — works precisely because of doc 07's prefix caching: the stable opening (system prompt, tools, skill index) is prefilled once at the cached rate, and each turn only prefills the delta. Compartmentalization keeps that prefix stable; stability keeps it cheap.
03 Why Harnesses Look Like This
The context window is the scarce resource — not disk, not network, not even model quality. Three pressures dictate the architecture:
04 How It Runs — One Harness Turn, Seven Steps
Step through a single orchestrator turn of a Claude Code–style harness and watch where tokens go — and where they don't.
The key accounting trick is step 6: the worker may have burned 30K tokens exploring a codebase, but the orchestrator's transcript only grows by the ~0.5K summary. That is doc 09's "compress via tool" pattern, realized as a process boundary: the orchestrator pays for complexity, workers pay for volume.
05 Context Assembly Math
Context assembly is a compiler: it takes source (memory bank, session state, skill files, transcript) and emits one token stream per request, optimizing for the cache. The emitted prompt has three regions:
Each region bills at a different rate (doc 07's table), so the cost of a turn is:
cost/request ≈ cached_prefix × 0.1 + fresh_input × 1.0 + output × 3
Worked example. Orchestrator: 12K stable prefix, 2K fresh per turn, 0.5K output per turn:
Growth rates. A naive single agent that pastes every tool output into one transcript accumulates T turns of unbounded context: per-turn cost grows linearly and total session cost grows O(T²) even with caching (the transcript itself is the growing prefix). Worse, attention quality decays as the transcript passes the effective window (doc 19). The subagent-constrained architecture bounds per-call context: orchestrator stays O(task complexity), each worker starts fresh with a tapered budget (≤ some threshold, e.g. 30–50K tokens), and peak context per call stays flat no matter how long the task runs.
06 Architecture Patterns Compared
Beyond the orchestrator–worker core, harnesses differ in who holds context and who decides. Anthropic's agent-building guidance names three archetypes — orchestrator–workers, planner–executor, and evaluator–optimizer — plus the decentralized extreme:
| Pattern | Context cost | Coordination overhead | When to use |
|---|---|---|---|
| Single agent + tools | One transcript; grows with total work — O(T²) session cost | Zero — but all exploration tax lands in one window | Short, focused tasks where tool outputs are small and summarizable |
| Orchestrator–worker (Claude Code–style) | Orchestrator O(task complexity); workers pay their own volume | N subagent calls; briefs must be written carefully | Parallelizable, read-heavy work: research, multi-file changes, fan-out search |
| Planner–executor | Planner context stays small (plan, not transcripts); executor per-step | Plan revision loop between the two roles | Long-horizon tasks where sequencing mistakes are expensive |
| Evaluator–optimizer | Two contexts: generator + critic; critic re-reads output each round | R rounds × re-evaluation cost | Tasks with a checkable quality bar: tests, specs, rubrics |
| Decentralized swarm | Many small contexts — minimal per-agent bloat | Highest: agents communicate peer-to-peer; no global view | Emerging/experimental; robustness over efficiency |
| Blackboard | Shared structured store; each specialist reads only its slice | Read/write policy design on the blackboard | Multi-expert problems with no fixed pipeline order (classic AI: HEARSAY-II) |
Planner–executor in one sentence
The planner never touches raw tool output; it produces and revises a plan artifact. Executors turn plan steps into actions. The plan is the interface — a natural compression point, since a plan is O(steps), not O(observations).
Blackboard in harness terms
A shared file or DB the orchestrator and workers both read/write — decisions, findings, open questions. It substitutes for shared context: agents coordinate through state on disk, which is cheap to store and selectively read, instead of tokens in a window.
07 The Bottleneck Mapping Table
The mapping discipline from doc 09, extended to the full harness. Every mechanism earns its complexity by relieving a specific bottleneck from docs 06–09:
| Mechanism | Bottleneck it relieves | Hardware/cost link |
|---|---|---|
| Cache-stable static prefix | Repeated prefill compute of identical opening tokens | 0.1× cached rate (doc 07); TTFT collapse |
| Subagent fresh contexts | Transcript growth taxing every later decode | KV-cache read per decode step (docs 06, 22) |
| Summaries as return values | Raw tool output entering the persistent transcript | Fresh-input + permanent attention tax |
| Tapered worker budgets | Unbounded single-context growth → dilution | Lost-in-the-middle (doc 19) |
| Memory bank (files/DB) | Window as working set, not database | Disk is ~10⁵ cheaper than context tokens |
| Skills: name+desc in context | Static knowledge inflating the cached prefix | Pays 0.1× on the index, 1.0× only on demand |
| RAG / hybrid retrieval | Relevance filtering before token spend | Embedding lookup ≪ prefill of wrong docs (doc 11) |
| Tool permission gate | Blast radius of injected instructions | Defense-in-depth (doc 16) |
| Write policies on memory | Write cost vs. read value asymmetry | Store decisions, not transcripts |
08 Memory, Skills, Retrieval — The Persistence Layer
Memory bank & procedural memory
Playbooks ("how we deploy here"), decisions, and lessons-learned live in files the harness reads into context when relevant. This is procedural memory: not what happened, but how to act. It survives session death because it lives on disk, priced in bytes, not tokens.
Episodic & comparative memory
What got persisted this session, and why. The write policy is an economics problem: raw transcripts have huge write cost and near-zero read value; decisions and lessons have tiny write cost and high read value. Persist the delta, not the log.
Skills = progressive disclosure, priced
A skill ships a name + one-line description in the cached prefix (a few dozen tokens each) and a detailed body on disk. The model sees the menu; the harness loads the dish only when the task matches. The cost math: 200 skills × ~30 tokens of description ≈ 6K tokens of menu at the 0.1× cached rate — versus megabytes of documentation that would otherwise be pasted, prefilled, and attention-taxed every turn. If a skill triggers, its body (k tokens) is paid at full rate once, on the turns that need it.
09 Failure Modes
Context rot & attention dilution
Long transcripts don't just cost more — they degrade. Mid-context information is systematically under-attended (doc 19). Symptom: the agent "forgets" an instruction stated 80K tokens ago even though it's still in the window. The fix is architectural: shorter effective contexts, summaries, re-injection of key constraints near the turn-local region.
Tool-call storms
A loop of failed searches or retry storms burns output tokens at 3–5× while adding junk to the transcript. Harnesses need budgets: max calls per task, loop detection, and hard termination — the control-loop layer's real job.
Prompt injection from tool output
Every tool result is untrusted input inside the context. A malicious web page read by a worker can command the orchestrator. Defenses are layered — permission gates, output quarantining, least-privilege tools (doc 16) — because context assembly cannot distinguish "text" from "instructions."
Coordination overhead blowup
Spawning N workers for a task one agent could do re-reads the stable prefix N times and serializes on brief-writing. Subagents are a compression bet; if the subtask's exploration volume is small, the bet loses. Measure: worker tokens ÷ summary tokens should be ≫ 1.
10 Mental Models
Context window = RAM (scarce, fast, priced per byte); memory bank = disk (cheap, persistent, needs explicit loads); subagents = processes (own address space, isolated, return exit codes); skills = dynamically linked libraries (mapped on demand); the control loop = the scheduler with a permission ring. Lets you reason about: why "the model" is the CPU and the harness is everything that makes a computer out of it.
Source files (memory, skills, plan) + symbol table (session state) → one emitted object file (the request). The cache is the compiler's incremental-build cache: touch nothing in the prefix and the build is nearly free. Lets you reason about: prompt ordering as a compilation constraint, not a style choice (doc 07).
A good orchestrator delegates: it writes a precise brief, gets a summary back, and makes the decision. It never reads the 30-page report — not because it can't, but because reading it would clog its one working memory forever. Lets you reason about: O(task complexity) vs. O(total work) contexts.
11 Anti-patterns
Bound the orchestrator's transcript with summaries and worker isolation; write decisions & lessons to the memory bank; keep the static prefix genuinely static; give workers tapered budgets and demand compressed returns.
Run the god-agent: one transcript, everything pasted in, "the window is 200K so who cares." That design pays O(T²) session cost, hits attention dilution, and rots — while still billing you at full price every turn.
Treat memory writes as investments with a read-value test: "will a future turn plausibly load this?" Persist the decision and the reason, not the transcript.
Dynamically inject volatile data into the cached prefix — timestamps, reordered tools, refreshed memory files mid-session. One changed token invalidates the entire suffix at 10× price (doc 07's fragility rule).
12 Misconceptions & Closing Insights
"More agents = better agent." Agents don't add intelligence; they add context isolation and parallelism — both of which cost coordination. A swarm that shares nothing reproduces work; an orchestrator with lazy briefs gets garbage summaries. Pattern choice follows the workload's volume-to-decision ratio.
"Subagents save tokens." They usually spend more total tokens (N fresh contexts, N prefix re-reads). What they save is the orchestrator's peak context — which is what actually compounds across turns and degrades quality. Cost-per-call down, sometimes total cost up; know which you're optimizing.
"Memory = a bigger context window." Long-context models raise the ceiling but not the economics: decode still reads the whole KV cache (docs 06, 22) and attention still dilutes (doc 19). The harness's memory hierarchy exists precisely because brute-force window growth is the wrong shape of solution.
"The harness is boilerplate around the model." Invert it: the model call is one opcode; the harness is the machine. Permissions, context assembly, memory policy, and worker scheduling determine cost, safety, and capability far more than the raw model choice at fixed quality tiers.
Related in this series
26 · AI Security Engineering · Perception · AI as an Observability Stack · 03 · Prompt & Context Engineering